ACS Catalysis
● American Chemical Society (ACS)
Preprints posted in the last 30 days, ranked by how well they match ACS Catalysis's content profile, based on 18 papers previously published here. The average preprint has a 0.01% match score for this journal, so anything above that is already an above-average fit.
Dorau, R.; Keller, M. B.; Thiesen, E. M.; Tiemann, J. K. S.; Gjermansen, M.; Tian, P.; Borch, K.; Jensen, K.; Westh, P.
Show abstract
Poly(ethylene terephthalate) (PET) is one of the most widely produced plastics, and enzymatic depolymerization offers a promising route to closed-loop recycling under mild conditions. However, most known bacterial PET hydrolases belong to a conserved canonical-fold cutinase family, leaving much of alpha/beta-hydrolase diversity unexplored. Here, we mapped bacterial cutinase sequence space by combining bioinformatics-guided sequence selection with high-throughput secretion screening in Bacillus subtilis. A library of 1,120 genes encoding 954 unique bacterial cutinases, spanning canonical- and minimal-fold families, was screened for activity on Impranil DLN and semicrystalline PET. We identified 156 secreted cutinases with polyester activity, broadly distributed across sequence space, but only ten showed detectable PET hydrolysis, all from the canonical-fold family. These PET hydrolases were active at 40-50{degrees}C, preferred alkaline pH, and showed moderate thermostability. Our results demonstrate that PET activity is rare among bacterial cutinases and provide a scalable workflow for discovering diverse enzyme starting points.
Ouyang, Y.; Nadeem, H.; Goto, Y.; Shukla, D.; van der Donk, W.
Show abstract
The biosynthetic machineries of ribosomally synthesized and post-translationally modified peptides (RiPPs) are often substrate tolerant. A remarkable example is the class II lanthipeptide synthetase ProcM, which naturally functions as a generalist enzyme that has not evolved to use a specific substrate during its evolutionary history. Although ProcM has been studied extensively, the sequence features associated with productive modification remain underexplored. In this study, we use the ultrahigh-throughput mRNA display technique to map the sequence compatibility of ProcM across a focused library. This approach expands the landscape of ProcM reactivity beyond native substrates and individually characterized variants. Machine learning (ML) is used as a tool to demonstrate that the selected dataset contains learnable signatures and classification architectures revealed a balanced accuracy of 0.73. This performance contrasts sharply with the near-perfect accuracy of specialized enzyme models as the sequence-fitness landscape of the generalist enzymes are characterized by class imbalance and limited by intrinsic dataset features. Our results provide a high-throughput view of ProcM reactivity and highlight differences with previous high-throughput studies on substrate selectivity of RiPP modification enzymes. Future studies will need to assess whether these differences are common when comparing generalist with specialist enzymes.
Nepogodiev, S.; Rejzek, M.; Steinberg, M. N.; Edwards, A.; Martin, C.
Show abstract
Oxalyl-coenzyme A (oxalyl-CoA) is a key intermediate in oxalate metabolism in plants, fungi and oxalate-degrading bacteria, but its limited availability has restricted biochemical investigations of oxalyl-CoA-dependent enzymes. Here, we describe a practical semisynthetic procedure for the preparation of oxalyl-CoA based on rapid oxalyl transfer from S-oxalyl p-thiocresol to coenzyme A. The reaction was monitored directly by 1H NMR spectroscopy, allowing optimisation of pD and reaction conditions. Following removal of thiocresol and purification by reversed-phase HPLC, oxalyl-CoA was obtained in 39% yield as determined by quantitative 1H NMR. The product was characterised by high-resolution electrospray mass spectrometry and comprehensive 1H, 13C and 31P NMR spectroscopy, confirming its structure unequivocally. During the study, the limited stability of oxalyl-CoA in aqueous solution was documented, leading to recommendations for its purification and storage. The semisynthetic protocol provides a convenient source of analytically pure oxalyl-CoA suitable for biochemical assays and supplies reference spectroscopic data for its unambiguous identification. The biological utility of the semisynthetic oxalyl-CoA was demonstrated by its application as an acyl donor substrate in assays of PnBAHD15, enabling quantitative kinetic characterisation of the enzyme and illustrating its suitability for biochemical studies of oxalyl-CoA-dependent enzymes.
Kobayashi, R.; Miyake, K.; Oya, T.; Ueno, H.; Saito, Y.; Noji, H.
Show abstract
The rotary motor F1-ATPase has been extensively studied as a model molecular machine, yet rational engineering of its catalytic activity remains challenging because ATP hydrolysis is regulated by long-range intersubunit allostery and large conformational transitions. Here, we developed a homolog-guided engineering strategy to increase the maximum rotation rate of the thermophilic Bacillus PS3 F1-ATPase (TF1). Candidate mutation sites were first identified by comparing TF1 with the homologous enzymes bovine mitochondrial F1 (bMF1) and Paracoccus denitrificans F1 (PdF1), both of which exhibit higher maximum rotation rates than TF1. Systematic exploration of these sites identified four activity-enhancing hotspots, followed by focused hotspot exploration and machine-learning-assisted prioritization of combinatorial mutants. The best mutant, TF1({beta}Y313L/{beta}E332S), exhibited a 1.8-fold higher maximum rotation rate than TF1(WT) while retaining its functional thermostability. Interestingly, activity-enhancing substitutions were not limited to the residues conserved in both bMF1 and PdF1, indicating that the bMF1-PdF1 consensus substitutions effectively identify activity-enhancing hotspots rather than uniquely defining the optimal amino acid. Machine-learning-assisted exploration efficiently prioritized highly active mutants, although the predictive performance was limited by the relatively small training dataset and epistatic interactions among mutations. Kinetic and structural comparisons further provided mechanistic insights into the enhanced catalytic activity of the engineered mutant. Together, these results establish a practical strategy for engineering complex molecular motors by combining homolog-guided hotspot identification with focused hotspot exploration.
Walkenhauer, E. G.; Cox-Tigre, N.; Chaubey, M.; Marcenac, R.; Wachsman, A.; Kodama, H. M.; Lindblom, K.; Bloom, C. E.; Antos, J. M.; Lisi, G. P.; Smirnov, S. L.; Amacher, J.
Show abstract
Bacterial sortase enzymes are cysteine transpeptidases at the surface of Gram-positive bacteria that ligate substrates to the cell wall. In addition, these enzymes are powerful tools in protein engineering applications via sortase-mediated ligation (SML) due to their covalent attachment of two substrates, with one containing a pentapeptide recognition motif with sequence LPXTG, where X=any amino acid, and the second, an N-terminal glycine. The class A sortase from Staphylococcus aureus (saSrtA) was the first to be identified, and over 25 years later, the most widely used SML variants continue to be derivatives of a directed-evolution-identified pentamutant of saSrtA, or saSrtA5M. We previously characterized P94, a position mutated in saSrtA5M that interacts directly with a structurally conserved loop (the {beta}7-{beta}8 loop) near the active site of wild-type saSrtA only in the inactive conformation. This work revealed that the single P94X mutation dramatically affects relative saSrtA activity, as well as specificity for the P2 (or X) position in the LPXTG recognition motif. This is largely driven by Km effects. Here, we further interrogated P94 by probing structural changes in the active, apo state of saSrtA in the presence of the P94D mutation, as well as via mutations in Y187, the {beta}7-{beta}8 loop residue hypothesized to interact directly with P94. The saSrtA enzyme is allosterically activated by calcium; therefore, we were interested if P94D would induce structural changes in the calcium-bound apo enzyme. We used 1H-15N NMR experiments to compare spectra between enzymatically inactive variants of saSrtA with and without the P94D mutation. We also used NMR to calculate relative binding affinities for a pentapeptide substrate to these variants, as well as enzymatically inactive saSrtA5M. Our NMR data, in combination with enzymatic assays using active variants confirmed differences in the active, apo states of these enzymes. Overall, this work provides additional atomic detail regarding the importance of the P94 residue in saSrtA substrate recognition.
Swartz, J.; Wang, W.; Liu, Q.
Show abstract
The declining cost of green hydrogen--projected below 1.5 USD/kg by 2030--opens new avenues for its use beyond fuel cells and industrial heating. Here we demonstrate that H2 can serve as a stoichiometric electron donor for cell-free enzymatic cofactor regeneration, coupling H2 oxidation to NADPH production and driving the complete bioconversion of pyruvate to lactate. A partially purified enzyme ensemble from Escherichia coli overexpressing Clostridium pasteurianum ferredoxin, augmented with [FeFe]-hydrogenase CpII, delivers NADP+ reduction rates of 103 M min-1 (27-fold enhancement) with superlinear dependence on H2 partial pressure. Reconstitution from purified components (CpI or CpII, CpFd, AnFNR, LDH) uncovers a redox-potential-dependent lag phase: the NADPH/NADP+ ratio must exceed 0.85 before pyruvate reduction becomes thermodynamically spontaneous, after which the rate accelerates exponentially. These results position hydrogen-driven cofactor regeneration as a scalable, byproduct-free platform for reductive biotransformations powered by renewable H2.
Furubayashi, M.
Show abstract
Nature produces hundreds of carotenoids, yet only a handful of the apocarotenoids derived from them are accessible through microbial production. The best-known example is retinal, the chromophore of rhodopsins and a precursor of pharmaceutical retinoids, which is generated by the central cleavage of {beta}-carotene. Whether the same cleavage chemistry can be extended to other carotenoids, yielding retinal analogues that differ in their ring structures, and potentially in their biological activities, has remained largely untested. In this study, we demonstrate a pathway engineering approach in E. coli for the biosynthesis of diverse retinal analogues by leveraging substrate promiscuity of Blh, a bacterial carotenoid cleavage enzyme originally identified in microbial rhodopsin gene clusters. While initial co-expression of Blh with carotenoid pathway genes often resulted in the production of retinal (by cleavage of {beta}-carotene intermediate), we found that by optimizing the expression level of Blh, carotenoids such as astaxanthin or canthaxanthin were cleaved efficiently. Structure-guided engineering of Blh, informed by its predicted substrate-binding cavity, further improved the cleavage of zeaxanthin. This expanded catalytic activity suggests that Blh can serve as a versatile biocatalyst for the production of diverse retinal analogues, potentially yielding compounds with a range of biological activities. Furthermore, our findings raise the possibility of diverse biological roles for these enzymes in their native biological contexts. ImportanceThis study demonstrated the successful biosynthesis of a diverse array of retinal analogues in engineered Escherichia coli through the heterologous expression of Blh, a {beta}-carotene cleavage dioxygenase, together with several carotenoid pathways. Careful design of the Blh expression construct enabled modulation of retinoid proportions in the engineered pathway. This work uncovers previously unrecognized substrate promiscuity of Blh, revealing its capacity to accept carotenoids beyond {beta}-carotene as substrates. For the first time, the predicted structure of Blh revealed the enzymes substrate cavity. Rational engineering by amino acid substitution designed to expand the cavity enabled the improved cleavage of hydroxylated carotenoids. These findings open new avenues for both fundamental research and biotechnological applications and have the potential to impact the microbial production of valuable retinoids.
Yasukochi, R.; Kashima, T.; Mori, T.; Kawauchi, Y.; Miyanaga, A.; Watanabe, H.; Fushinobu, S.
Show abstract
Cyclic oligosaccharides possess industrial advantages, including molecular encapsulation capability and high physicochemical stability, owing to the absence of a reducing end. Recently, a novel cyclic tetrasaccharide, cycloisomaltotetraose (CI4), consisting of four -1,6-linked glucose units, and the enzymes responsible for its synthesis, cycloisomaltotetraose glucanotransferases (CI4Tases), were discovered. Unlike known cycloisomaltooligosaccharide glucanotransferases (CITases) that yield a wide distribution of cyclic products with a degree of polymerization (DP) of 7 or higher, CI4Tases strictly produce CI4. To elucidate the molecular mechanism underlying this strict DP4 specificity, we determined the crystal structures of CI4Tase from Agreia sp. D1110, in its ligand-free form, as well as in complex with the linear hydrolysis product isomaltotetraose (IG4) and with CI4. Structural comparisons revealed that a loop (M247 to R251) blocks the region corresponding to the -5 subsite of typical CITases, narrowing the substrate-binding pocket. This "molecular ruler" mechanism ensures that only a glycan chain of exactly four glucose units is accommodated for cyclization. Among mutants of the residue positioned at the center of bound CI4, the formation of by-products other than CI4 was significantly suppressed in F245L, F245A, and F245W. While the cyclization activity of all F245 mutants decreased, the CI4 hydrolysis activity of these three mutants was also significantly reduced, resulting in an increased specificity for cyclic sugar production. These findings elucidate the strict size-control mechanism of CI4Tase and provide a structural foundation for engineering cycloisomaltooligosaccharide-producing enzymes with optimized transglycosylation efficiency and specificity for industrial applications.
Tomlinson, C. W.; Elli, S.; Batiste-Simms, M.; Chen, Z.; Taylor, C.; Dowle, A.; Yates, E. A.; Nazare, M.; Fascione, M.; Willems, L.; Williams, S. J.; Crawford, C. J.; Cartmell, A.
Show abstract
The enzymatic removal of sulfate groups regulates processes ranging from steroid metabolism to carbohydrate degradation. Most sulfatases belong to the S1 family, whose members use a co-translationally installed formylglycine residue to hydrolyse sulfate esters. Arylsulfamates are potent covalent inhibitors of aryl and steroid sulfatases, including the clinical steroid sulfatase inhibitor Irosustat, yet the structure and stability of the inhibited complex remain unresolved. Arylsulfamates and carbohydrate sulfamates do not covalently inhibit many S1 carbohydrate sulfatases despite conservation of their sulfate-binding sites and formylglycine residue. Using enzyme kinetics, X-ray crystallography, molecular dynamics simulations and density functional theory calculations, we define the basis of these contrasting behaviours. High-resolution structures of the Pseudomonas aeruginosa arylsulfatase PaAtsA treated with two arylsulfamates reveal a long-lived tetrahedral, O-linked -hydroxysulfamate adduct attached to formylglycine. Molecular simulations show that replacing sulfate with sulfamate disrupts the favourable Ca2+-oxyanion interaction and alters ligand binding geometry. The permissive hydrophobic binding site of PaAtsA accommodates this rearrangement while retaining a trajectory compatible with nucleophilic attack. By contrast, in the Bacteroides thetaiotaomicron carbohydrate sulfatase BT16363S-Gal, sulfate-to-sulfamate substitution weakens binding and displaces the sulfamate from a reactive pose near the catalytic nucleophile due to a restrictive active site with conserved sugar binding. These findings define the structure and persistence of the arylsulfamate-derived covalent intermediate and explain why sulfamate warheads are tolerated by aryl sulfatases but not carbohydrate sulfatases.
Wang, H.; Mai, B. K.; Zhang, X.; Li, C.; Liu, P.; Yang, Y.
Show abstract
The cooperative integration of photoredox catalysis and metalloenzyme catalysis has emerged as a powerful strategy for enabling stereoselective radical transformations beyond the capabilities of either catalytic mode alone. Herein, we report a photometallobiocatalytic enantioselective intermolecular C-C cross-coupling of pyridotriazoles and secondary alkyltrifluoroborate salts through cooperative catalysis between an organic photosensitizer and an engineered protoglobin. By combining visible-light-mediated radical generation with enzymatic activation of pyridotriazoles to form reactive Fe carbenoid intermediates, this transformation enabled highly enantioselective radical C-C bond formation through a proposed outer-sphere coupling mechanism. Through biocatalyst mining and directed evolution, engineered Aeropyrum pernix protoglobin catalysts were developed that catalyzed this radical C-C coupling with excellent efficiency and stereocontrol. The photobiocatalytic platform exhibited a broad substrate scope with respect to both secondary alkyltrifluoroborate salts and pyridotriazoles, affording a range of valuable N-heterocyclic products in excellent yields and enantioselectivities. Mechanistic studies supported the involvement of radical intermediates and revealed spontaneous binding between the photocatalyst eosin B and the engineered metalloenzyme. By leveraging cooperative photometallobiocatalysis, this work established an underexplored strategy for asymmetric intermolecular radical cross-coupling via an outer-sphere mechanism, further expanding the catalytic repertoire of transition-metal carbenoid chemistry. Entry for the Table of Contents O_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=83 SRC="FIGDIR/small/744224v1_ufig1.gif" ALT="Figure 1"> View larger version (12K): org.highwire.dtl.DTLVardef@132b69corg.highwire.dtl.DTLVardef@72eea5org.highwire.dtl.DTLVardef@1919e26org.highwire.dtl.DTLVardef@125fba6_HPS_FORMAT_FIGEXP M_FIG An enantioselective photometallobiocatalytic cross-coupling of pyridotriazoles and secondary alkyltrifluoroborate salts is developed. Cooperative catalysis using eosin B and an engineered protoglobin combines visible-light-mediated radical generation with enzymatic metal carbenoid activation, affording valuable N-heterocyclic products in excellent yield and enantioselectivity through an outer-sphere radical coupling pathway. C_FIG
Batey, R. T.; Olenginski, L. T.; Wierzba, A. J.; Patel, D.
Show abstract
Contemporary RNA-binding ligand collections are biased toward aromatic scaffolds, although it remains unclear whether this over-representation reflects an intrinsic requirement for productive RNA recognition or historical discovery bias. Here, using a modular "host-guest" ligand design strategy targeting the env8 cobalamin (Cbl) riboswitch, we established a common molecular framework to directly evaluate whether aromaticity is fundamentally required for RNA binding. We synthesized a focused series of cyclic aliphatic {beta}-axial Cbl derivatives, expanding the ligand library and enabling matched-pair comparisons to isolate the contribution of aromaticity to molecular recognition. Aliphatic ligands supported high-affinity RNA binding and regulatory activity comparable to aromatic analogues, with several derivatives exhibiting equal or greater affinity than their matched aromatic counterparts. Structural analyses revealed that aromatic and aliphatic ligands engage the same cryptic RNA binding site through distinct modes of molecular recognition, including nucleobase {pi}-stacking and alternative van der Waals packing arrangements. Machine learning analyses further demonstrated that the physicochemical features associated with affinity extend beyond aromaticity itself and instead reflect a broader combination of shape, surface, heteroatom, and electronic properties. Together, these findings demonstrate that high-affinity RNA binding can arise from multiple structural and physicochemical solutions, suggesting that aromaticity is not uniquely privileged as a strategy for RNA-targeted ligand design and supporting broader exploration of underrepresented RNA-binding chemotypes.
Ruta, G. V.; Ciciani, M.; De Sanctis, V.; Bertorelli, R.; Valentini, C.; Menghini, D.; Kheir, E.; Gentile, M. D.; Conci, A.; Casini, A.; Cereseto, A.
Show abstract
Compact Cas nucleases offer advantages over the widely used SpCas9 due to their smaller size, which enables more efficient delivery for in vivo applications. Among these, the phage-encoded Cas{Phi}2 (Cas12j2) is highly promising due to its relaxed PAM requirement (5-TTN-3) and compact size (757 aa); however, its translational potential is limited by low editing activity. To enhance the efficacy of Cas{Phi}2, we optimized the previously reported EPICA system, developing EPICA.2, a eukaryotic directed evolution platform to improve nucleases with nearly undetectable activity. EPICA.2 integrates additional yeast evolution rounds to enrich for active variants along with a low background mammalian reporter system that improves detection and selection of enhanced variants. Finally, we set up a long-read sequencing protocol which uses unique molecular identifiers (UMIs) to reduce sequencing errors, enabling accurate identification of the mutation combinations in each evolved variant. Among the most frequent variants, we obtained evoCas{Phi}2, which contains six activity-boosting mutations with a synergistic effect not predictable by rational engineering. Overall, evoCas{Phi}2 showed up to 70-fold increased activity in human cells compared to wild-type and outperformed variants generated through rational approaches, highlighting the potential of EPICA.2 as a powerful strategy to evolve genome editing tools with low native activity.
Munnoch, J. T.; Larcombe, D. E.; McHugh, R. E.; Bruce, J.; Robb, K.; Croxford, J. T.; Kiepas, A. B.; Gomez-Escribano, J. P.; Crowhurst, N. A.; Collis, A. J.; Kendrew, S. G.; Huckle, B. D.; Wilkinson, B.; Hunter, I. S.; Hoskisson, P. A.
Show abstract
The domestication of Streptomyces species for antibiotic production involves long-term, iterative mutagenesis and selection, yet the genomic changes driving enhanced production remain unclear. Analysis of five strains from an industrial lineage of Streptomyces clavuligerus using comparative genomics, transcriptomics and phenotypic profiling for dynamic genome architectures with plasmid integration events and chromosomal rearrangements, alongside the accumulation of mutations affecting metabolic pathways and global gene regulation. These changes increased precursor supply and reprogrammed transcription leading to enhanced clavulanic acid production but reduced catabolic flexibility. Complementation experiments confirmed the functional impacts of specific mutations. These findings reveal that artificial selection shapes genome evolution in industrial strains, balancing production gains with metabolic trade-offs. This work will likely inform rational design of Streptomyces strains for improved natural product production in industry while highlighting the constraints imposed by domestication on metabolic versatility. More broadly it shows that many of the evolutionary processes in industrial strain improvement programmes mirror those at play during natural selection.
Navaratna, T. A.; Akram, J.; Pazdernik, T. D.; Ramachandran, A.; Schultz, P.; Dulchavsky, M.; Choussat, X.; Oczon, C.; Singh, A.; Myers, N.; Robida, A.; Tripathi, A.; Stull, F.; Bardwell, J. C.
Show abstract
NicA2 is a flavin-bound amine dehydrogenase from Pseudomonas putida S16 that converts nicotine to the pharmacologically inactive N-methylmyosmine. In animal models of nicotine addiction, injection of NicA2 can decrease nicotine-seeking behavior 10-fold. Accordingly, NicA2-related enzymes have been investigated as smoking-cessation therapeutics. However, efficient catalysis by NicA2 in Pseudomonas putida relies on electron transfer to CycN, a cytochrome c, and not directly to O2. Impractically high amounts of NicA2 are thus necessary to achieve a pharmacological effect in the absence of CycN. Directed evolution has improved the ambient-O2 value of kcat from 0.007 s-1 to 1 s-1 for NicA2, but further improvements have been challenging. Here, we identify a strain of Peribacillus frigoritolerans NIC8 which encodes two flavin amine oxidoreductases, Ncox and Pnox. In the presence of oxygen, Ncox and Pnox act on nicotine and pseudooxynicotine respectively with apparent kcat values of 7.7 s-1 and 3.9 s-1. Transient kinetics establishes bimolecular rate constants of 51100 M-1s-1 and 81000 M-1s-1 for the half-reactions between Ncox and O2 and between Pnox and O2 respectively, consistent with Ncox and Pnox being bona-fide oxidases. Transcriptomics shows enhanced expression of Ncox and Pnox under nicotine-dependent growth as well as supporting the identification of downstream enzymes. Phylogenetic analysis suggests that Ncox and Pnox arose out of repurposing of homologous enzymes found in Bacillus species. The enzymes we describe may be useful for the development of nicotine addiction therapeutics and for bioconversion of nicotine in waste streams.
Sapienza, P. J.; Vera-Rodriguez, D. J.; Mileur, T. R.; Lee, A. L.
Show abstract
The classical understanding of allostery was initially grounded in two-state models, such as MWC and KNF, where structure and function are inextricably linked through transitions between low-(T) and high-affinity (R) states. Here, we show Yeast chorismate mutase (CM) provides a vivid example of the growing list of exceptions to the traditional T vs R two-state allosteric paradigm. While CM exhibits dynamic sampling of the R-state in the presence of the activator tryptophan (Trp), suggesting a conformational selection (CS) mechanism, we present multiple instances where conformational status and catalytic activity are decoupled. Using NMR spectroscopy and kinetic assays, we identify CM variants that reside almost exclusively in the T conformation can exhibit maximal activity, while others that predominantly occupy the R conformation are weakly active. Quantitative comparison of experimental data with a parameterized CS model reveals deviations of up to two orders of magnitude, ruling out the simplest two-state model for substrate affinity modulation in this system. We propose that the observed T-to-R switching in CM is "incidental", a byproduct of an evolved energy landscape that allows access to the substrate-bound pose but does not mechanistically determine affinity. Our findings suggest that allosteric regulation in CM may instead be driven by local features of the ground-state ensemble, which operate independently of global T/R status. This work further highlights an emerging view that the mere observation of a pre-sampled active conformation does not sufficiently prove a two-state mechanism and further underscores the need for deeper ensemble-based perspectives in protein engineering and allostery.
Toneyan, S.; Scholz, K.; De Donno, C.; Noack, F.; Auslaender, S.; Cijsouw, T.; Payne, J. L.
Show abstract
Codon optimization uses synonymous sequence changes to improve the expression and therapeutic performance of nucleic acid-based medicines. Masked language models (MLMs) have recently been proposed as alternatives to traditional, frequency-based codon optimization approaches, yet whether they offer a meaningful advantage over such simpler methods remains unclear. Here we benchmark three prominent MLMs - CaLM, EnCodon and CodonTransformer - across backtranslation fidelity, sequence generation and nine molecular phenotype prediction tasks, and experimentally evaluate model-designed sequences using a secreted embryonic alkaline phosphatase (SEAP) reporter. The models differed markedly in amino-acid fidelity and generated distinct synonymous sequence variants. However, no single model performed best across all benchmark tasks and simple sequence features remained competitive in several settings. Our interpretability analysis revealed that the models integrate a large window of codon context for making predictions, as opposed to frequency-based approaches. Our in vitro data showed that MLM-designed variants outperformed conventional and commercial-vendor-derived sequences in both transient and stably integrated expression, supporting the models ability to capture translational context beyond codon frequency. Together, our results establish MLMs as effective and complementary tools for codon optimization and suggest that sampling across multiple models may improve the likelihood of identifying high-performing therapeutic sequences.
Jankovicova, B.; Bigos, A.; Surpeta, B.; Silva, M.; Brezovsky, J.; Dvorak, P.
Show abstract
Efficient conversion of polymeric feedstocks for sustainable bioprocessing requires robust strategies for enzyme assembly and cell-surface attachment. In nature, cellulosomes achieve highly efficient lignocellulosic polysaccharide deconstruction through scaffoldin-mediated organization of carbohydrate-active enzymes via specific cohesin-dockerin interactions. These modular binding pairs are therefore attractive tools for synthetic biology and engineered whole-cell biocatalysis, yet their performance has been studied mainly in vitro or in yeast or Gram-positive bacteria. The factors governing their function on the microbial surfaces - particularly those of Gram-negative bacteria - remain incompletely understood. Here, we investigated the binding efficiency and interaction stability of two thermophilic cohesin-dockerin pairs from Acetivibrio thermocellus and Acetivibrio clariflavus displayed on the surface of the genome-streamlined strain Pseudomonas putida EM371 using an Ag43-based display system from Escherichia coli and a dockerin-tagged fluorescent reporter. We show that binding efficiency is strongly affected by the temperature at which the cohesin-dockerin complex is formed. We further demonstrate that the interaction stability of the A. clariflavus pair can be substantially improved by targeted amino acid substitutions in the dockerin domain guided by molecular dynamics simulations and free-energy calculations. These results identify key parameters controlling the performance of thermophilic cohesin-dockerin modules on living bacterial cell surfaces and establish a computation-guided strategy for engineering more stable cellulosome-derived assembly interfaces, advancing the development of modular whole-cell platforms for sustainable biotechnology applications. TOC graphics O_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=107 SRC="FIGDIR/small/743725v1_ufig1.gif" ALT="Figure 1"> View larger version (65K): org.highwire.dtl.DTLVardef@18a97b6org.highwire.dtl.DTLVardef@1ee3ff4org.highwire.dtl.DTLVardef@a8dd60org.highwire.dtl.DTLVardef@5dd332_HPS_FORMAT_FIGEXP M_FIG C_FIG Cohesin-dockerin pairs provide strong and modular non-covalent interactions for synthetic biology and biotechnology applications. We establish an experimental and computational pipeline to improve their two key properties - binding efficiency and interaction stability - on the surface of Pseudomonas putida, enabling more robust cell-surface assembly systems.
Chen, K.; Qi, Z.; Lozano Ramos, O.; Li, H.; Ma, M.; Gannarapu, M. R.; Bi, F.; Li, A.; Li, H.; XIONG, R.
Show abstract
AlphaFold 3 (AF3) and Boltz-2 are state-of-the-art AI-based tools for biomolecular structure prediction, but whether their predictions provide useful guidance for lead optimization, SAR interpretation, and virtual screening remains insufficiently characterized. We benchmarked their performance using newly determined soluble epoxide hydrolase co-crystal structures and matched activity data together with a curated post-training-cutoff dataset spanning kinases, allosteric modulators, covalent systems, PROTACs, molecular glues, fragments, membrane proteins, RNA binders, and activity-cliff pairs. Both models recovered canonical orthosteric enzyme and kinase complexes, including key DFG/C conformational states, whereas allosteric, membrane-protein, and induced-proximity complexes remained challenging. Pharmacophore RMSD was often lower than overall ligand RMSD, indicating preservation of key recognition features despite imperfect whole-ligand alignment. AF3 minPAE correlated with pose accuracy, and very low minPAE values (<0.85 A) were strongly enriched for accurate poses. Model confidence scores were not associated with experimental activity, whereas Boltz-2 predicted affinity captured relative activity trends and distinguished the activity-cliff pair, although its performance varied across ligand series.
Ouchida, S. T.; Horst, M. T.; Gou, X.; Bakanas, I.; Hatstat, A. K.; Schnaider, L.; Diolaiti, M. E.; Ashworth, A.; DeGrado, W. F.
Show abstract
The de novo design of proteins that bind chemically complex small molecules has broad chemical and biological implications, but strategies typically rely on a small set of protein scaffolds and require extensive experimental screening. Here, we computationally designed proteins around a minimal aromatic {pi}-stacking motif to bind the anthracycline anticancer drug doxorubicin. Experimental characterization of twelve proteins revealed a {micro}M doxorubicin binder; two additional design cycles improved scaffold stability and binding affinity to yield an 85-residue protein that binds doxorubicin with a dissociation constant of 85 nM. An X-ray crystal structure of the protein-drug complex confirmed the accuracy of the designed {pi}-{pi} stacking interactions. The designed protein could act to protect cultured cells from doxorubicin-induced cytotoxicity. Unlike previous ligand-binding protein designs based on repeat proteins or naturally occurring folds, the designed protein adopts a previously unobserved 5-helix globular fold, indicating that a broader space of folded, functional proteins exists even for compact tertiary structures smaller than 100 residues. These results demonstrate that motif-guided generative protein design can discover compact de novo protein folds capable of high-affinity recognition of chemically complex small molecules.
Nishioka, R.; Murozono, K.; Kawaguchi, Y.; Kimura, M.; Sakuraba, S.; Hashii, N.; Senoo, A.; Caaveiro, J.; Umetsu, M.; Kamiya, N.
Show abstract
Site-specific protein modification allows diverse functionalities to be introduced while minimizing perturbations to the protein structure and activity. Considerable efforts have been made to achieve site-specific modification of native proteins to overcome the heterogeneity resulting from conventional stochastic Lys or Cys modification. We have previously achieved the selective modification of Lys65 in a native immunoglobulin G1 (IgG1) antibody (trastuzumab) using EzMTG-pG(Fab), which is an engineered zymogen of microbial transglutaminase (EzMTG) fused to a Fab-binding protein G [pG(Fab)]. However, this approach cannot be widely applied to different types of IgG antibodies. Here, we designed pG(Fab)-EzMTG by fusing pG(Fab) to the N-terminus of EzMTG. Notably, switching the fusion partners dramatically altered the IgG modification site from Lys65 to Lys225, which is located in the hinge site of native IgG1 antibodies. This Lys225-selective labeling was applicable to different IgG1 antibodies. As a functional application, the cytotoxic drug monomethyl auristatin E (MMAE) was conjugated to Lys225 of trastuzumab, and the resulting antibody-drug conjugate exhibited antigen-specific cytotoxicity. These findings demonstrate that fusion-protein architecture determines site selectivity in proximity-directed enzymatic modification, providing a strategy for the site-specific functionalization of native antibodies.